NeedleBench: Can LLMs Do Retrieval and Reasoning in Information-Dense Context?

Li, Mo; Zhang, Songyang; Zhang, Taolin; Duan, Haodong; Liu, Yunxin; Chen, Kai

Computer Science > Computation and Language

arXiv:2407.11963 (cs)

[Submitted on 16 Jul 2024 (v1), last revised 9 May 2025 (this version, v2)]

Title:NeedleBench: Can LLMs Do Retrieval and Reasoning in Information-Dense Context?

Authors:Mo Li, Songyang Zhang, Taolin Zhang, Haodong Duan, Yunxin Liu, Kai Chen

View PDF HTML (experimental)

Abstract:The capability of large language models to handle long-context information is crucial across various real-world applications. Existing evaluation methods often rely either on real-world long texts, making it difficult to exclude the influence of models' inherent knowledge, or introduce irrelevant filler content to artificially achieve target lengths, reducing assessment effectiveness. To address these limitations, we introduce NeedleBench, a synthetic framework for assessing retrieval and reasoning performance in bilingual long-context tasks with adaptive context lengths. NeedleBench systematically embeds key data points at varying depths to rigorously test model capabilities. Tasks are categorized into two scenarios: information-sparse, featuring minimal relevant details within extensive irrelevant text to simulate simple retrieval tasks; and information-dense (the Ancestral Trace Challenge), where relevant information is continuously distributed throughout the context to simulate complex reasoning tasks. Our experiments reveal that although recent reasoning models like Deepseek-R1 and OpenAI's o3 excel in mathematical reasoning, they struggle with continuous retrieval and reasoning in information-dense scenarios, even at shorter context lengths. We also characterize a phenomenon termed 'under-thinking', where models prematurely conclude reasoning despite available information. NeedleBench thus provides critical insights and targeted tools essential for evaluating and improving LLMs' long-context capabilities. All resources are available at OpenCompass: this https URL.

Comments:	v2: updated with tested models and Multi-Needle Reasoning implementation
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2407.11963 [cs.CL]
	(or arXiv:2407.11963v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2407.11963

Submission history

From: Mo Li [view email]
[v1] Tue, 16 Jul 2024 17:59:06 UTC (1,092 KB)
[v2] Fri, 9 May 2025 09:23:22 UTC (1,257 KB)

Computer Science > Computation and Language

Title:NeedleBench: Can LLMs Do Retrieval and Reasoning in Information-Dense Context?

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:NeedleBench: Can LLMs Do Retrieval and Reasoning in Information-Dense Context?

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators